Papers with embedding layer
Language Directions in Multilingual LLMs: A Layer-wise Diagnostic Study of Token Alignment and Pretraining Imprint (2026.acl-srw)
Copied to clipboard
| Challenge: | Using a unified probing framework, we analyze six multilingual LLMs across five languages. |
| Approach: | They analyze multilingual representations across five languages and analyze their behavior . they find that accuracy rises by +73.5 to +80.7 points from L0 to L1 on average . |
| Outcome: | The proposed framework enables a consistent and substantial early jump in accuracy across models . the token–language alignment measures where vocabulary sharing peaks . |
Neural Machine Translation without Embeddings (2021.naacl-main)
Copied to clipboard
| Challenge: | Existing models operate over subword tokens, but byte-based models employ a different approach . a one-hot representation of each byte does not hurt performance, but it improves BLEU scores . |
| Approach: | They propose to represent every computerized text as a sequence of bytes via UTF-8 . this eliminates the need for an embedding layer and improves performance . |
| Outcome: | The proposed model improves BLEU scores on byte-to-byte translation models compared to character-level models . the proposed model does not require an embedding layer and does not drop out of the decoder . |
KGLM: Integrating Knowledge Graph Structure in Language Models for Link Prediction (2023.starsem-1)
Copied to clipboard
| Challenge: | Knowledge graphs are incomplete in the information they represent, necessitating knowledge graph completion tasks. |
| Approach: | They propose a new entity/relation embedding layer that learns to differentiate distinctive entity and relation types, thus allowing the model to learn the structure of the knowledge graph. |
| Outcome: | The proposed language model learns to differentiate distinct entity and relation types, thus learning the structure of the knowledge graph. |
Towards Simple and Efficient Task-Adaptive Pre-training for Text Classification (2022.aacl-short)
Copied to clipboard
| Challenge: | Large-scale pre-trained language models are extensively trained on massive heterogeneous datasets, known as pre-training datasets. |
| Approach: | They propose to use Domain Adaptive Pre-training and Task-Adaptive pre-training as intermediate steps before the final finetuning task to cover the target domain vocabulary. |
| Outcome: | The proposed approach is computationally efficient, with 78% fewer parameters trained during TAPT. |
Improving Neural Machine Translation by Incorporating Hierarchical Subword Features (C18-1)
Copied to clipboard
| Challenge: | Using subwords, we find that the appropriate subword units for the three layers differ depending on the model . incorporating hierarchical subword features improves BLEU scores on the IWSLT evaluation datasets. |
| Approach: | They propose a method that expresses a word by combining "subwords" they propose to incorporate hierarchical subword features into a single embedding layer . |
| Outcome: | The proposed method improves BLEU scores on the IWSLT evaluation datasets. |
Multimodal Pretraining Unmasked: A Meta-Analysis and a Unified Framework of Vision-and-Language BERTs (2021.tacl-1)
Copied to clipboard
| Challenge: | Large-scale pretraining and task-specific fine-tuning are now the standard methodology for many tasks in computer vision and natural language processing. |
| Approach: | They propose to combine two types of vision and language BERTs to create a theoretical framework that can be unified under different theoretical frameworks. |
| Outcome: | The proposed models can be classified into single-stream or dual-stream encoders and are unified under a single theoretical framework. |
Efficient Vocabulary Reduction for Small Language Models (2025.coling-industry)
Copied to clipboard
| Challenge: | Large language models (LLMs) have high computational costs and energy consumption, making their deployment in industrial settings difficult. |
| Approach: | They propose a small language model that compresses the embedding layer and reduces model size without significant loss of performance. |
| Outcome: | The proposed model reduces the embedding layer while maintaining performance while improving accuracy and performance. |
WSpeller: Robust Word Segmentation for Enhancing Chinese Spelling Check (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Chinese spelling check (CSC) detects and corrects spelling errors in Chinese texts. |
| Approach: | They propose a Chinese spelling check model that takes into account word segmentation and a module that can assist the correction module by predicting correct word segmentations from sentences containing spelling errors. |
| Outcome: | The proposed model outperforms baselines on SIGHAN13, SIGHEN14, and SIGHAN15 and maintains equal performance on SSGHAN14. |
KroneckerBERT: Significant Compression of Pre-trained Language Models Through Kronecker Decomposition and Knowledge Distillation (2022.naacl-main)
Copied to clipboard
| Challenge: | a recent study shows that over-parameterized pre-trained language models are unsuitable for low-capacity devices. |
| Approach: | They propose a transformer-based pre-trained language model that is overparameterized . they use a two-stage knowledge distillation scheme to train the model . |
| Outcome: | The proposed model outperforms state-of-the-art models on well-known NLP benchmarks. |
Gradient Inversion Attack in Federated Learning: Exposing Text Data through Discrete Optimization (2025.coling-main)
Copied to clipboard
| Challenge: | federated learning could overcome the bottleneck of public text data in large language models . a novel attack method is proposed to fully expose text data from gradients . |
| Approach: | They propose a method to fully expose text data from gradients by using a network of clients and a server. |
| Outcome: | The proposed method shows it is possible to Fully Expose Text data from gradients. |
On Romanization for Model Transfer Between Scripts in Neural Machine Translation (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Using romanization to improve low-resource machine translation is not always the best strategy. |
| Approach: | They propose to use romanization to improve transfer between languages with different scripts . they compare two romanization tools and find that they exhibit different degrees of information loss, which affects translation quality. |
| Outcome: | The proposed method improves transfer between languages with different scripts while entails information loss. |
Game-theoretic Vocabulary Selection via the Shapley Value and Banzhaf Index (2021.naacl-main)
Copied to clipboard
| Challenge: | Using the full vocabulary results in less explainable and memory intensive models. |
| Approach: | They propose a vocabulary selection method that views words as members of a team trying to maximize the model's performance. |
| Outcome: | The proposed method outperforms baseline models on multiple tasks and datasets. |
Wav-BERT: Cooperative Acoustic and Linguistic Representation Learning for Low-Resource Speech Recognition (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods to learn the transfer from speech to text are unexplored . how to solve the representation discrepancy of speech and text is unexplorable . |
| Approach: | They propose a cooperative acoustic and linguistic representation learning method to fuse and utilize contextual information of speech and text. |
| Outcome: | The proposed method outperforms existing methods on low-resource speech recognition. |
Bayesian Compression for Natural Language Processing (D18-1)
Copied to clipboard
| Challenge: | In natural language processing, recurrent neural networks have a huge number of parameters. |
| Approach: | They propose a Bayesian sparsification technique which allows compressing RNNs dozens or hundreds of times without time-consuming hyperparameters tuning. |
| Outcome: | The proposed technique compresses the RNN dozens or hundreds of times without time-consuming hyperparameters tuning. |
Tokenizer-Aware Cross-Lingual Adaptation of Decoder-Only LLMs through Embedding Relearning and Swapping (2026.eacl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have been primarily focused on English, leaving the multilingual ability unexplored. |
| Approach: | They propose a technique that creates new tokenizers and tunes embeddings on fixed model weights for target language adaptation. |
| Outcome: | The proposed method is light-weight and performant but has limitations for older models and high resource languages. |
Adapting Pre-trained Language Models to African Languages via Multilingual Adaptive Fine-Tuning (2022.coling-1)
Copied to clipboard
| Challenge: | Multilingual pre-trained language models have shown impressive performance on several downstream tasks for both high-resourced and low-resource languages. |
| Approach: | They propose to apply multilingual adaptive fine-tuning to 17 most-resourced African languages and three other high-resource languages to encourage cross-lingual transfer learning. |
| Outcome: | The proposed approach is competitive to LAFT on individual languages while requiring significantly less disk space. |
Bridge the Gap Between CV and NLP! A Gradient-based Textual Adversarial Attack Framework (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for adversarial samples are poorly applied in computer vision . however, textual adversarials are still vulnerable to small perturbations . |
| Approach: | They propose a framework to extend existing adversarial attack methods to textual adversarials by adding optimized perturbations to embedding layer and amplifying them in forward propagation process. |
| Outcome: | The proposed framework achieves better performance even using proxy gradient information and produces more fluent and grammatical adversarial samples compared to baseline methods. |
Linguistically Informed Hindi-English Neural Machine Translation (2020.lrec-1)
Copied to clipboard
| Challenge: | Neural Machine Translation (NMT) is a promising approach to machine translation . lack of parallel training data for Hindi-English is limiting . |
| Approach: | They propose to incorporate linguistic knowledge encoded by Hindi phenomena into a Transformer model to improve the translation performance. |
| Outcome: | The proposed model incorporates linguistic features to improve the translation performance. |
KS-Lottery: Finding Certified Lottery Tickets for Multilingual Transfer in Large Language Models (2025.naacl-long)
Copied to clipboard
| Challenge: | Existing studies have shown that a small subset of parameters is highly effective in fine-tuning . prior work shows that there are a few additional parameters corresponding to an intrinsic dimension in a well-trained Large Language Model. |
| Approach: | They propose a method to identify a small subset of LLM parameters highly effective in multilingual fine-tuning. |
| Outcome: | The proposed method can find the certified winning tickets in the embedding layer, and fine-tuning on the found parameters is guaranteed to perform as well as full fine- tuning. |
Gamma-Guard: Lightweight Residual Adapters for Robust Guardrails in Large Language Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large language models (LLMs) are widely deployed as zero-shot evaluators for answer grading, content moderation, and document ranking. |
| Approach: | They propose a system that trains LLMs with adapters to denoise embeddings and refocus attention. |
| Outcome: | The proposed model lifts adversarial accuracy from 5% to 95% a 90 percentage-point gain while reducing clean-data accuracy by just 8 percentage points. |
CoVE: Compressed Vocabulary Expansion Makes Better LLM-based Recommender Systems (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing approaches to align LLMs with recommendation tasks do not fully leverage their sequential information processing capabilities. |
| Approach: | They propose a system that allows users to expand their vocabulary by assigning a unique ID to each item within the expanded vocabulary. |
| Outcome: | The proposed system maximizes the sequence understanding abilities of large language models, significantly enhancing their performance on recommendation tasks. |
Structured Pruning for Efficient Generative Pre-trained Language Models (2023.findings-acl)
Copied to clipboard
| Challenge: | Large-scale generative Pre-trained Language Models (PLMs) are limited in their deployment in real-world applications. |
| Approach: | They propose to prune the feed-forward networks of generative pre-trained language models to smaller widths without designing extra operators. |
| Outcome: | The proposed method achieves 1.51x/6.96x inference speedup on GPU/CPU with 67% size reduction. |
Spelling-out is not Straightforward: LLMs’ Capability of Tokenization from Token to Characters (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models (LLMs) can spell out tokens character by character with high accuracy, yet struggle with more complex character-level tasks. |
| Approach: | They examine how large language models internally represent character-level information during the spelling-out process. |
| Outcome: | The embedding layer does not fully encode character-level information, especially beyond the first character. |
Attention-Focused Adversarial Training for Robust Temporal Reasoning (2022.lrec-1)
Copied to clipboard
| Challenge: | Current adversarial training approaches for NLP add adversarials to the embedding layer, ignoring other layers. |
| Approach: | They propose an enhanced adversarial training algorithm for fine-tuning transformer-based language models . they add the adversarials to multiple hidden states or attention representations of the model layers . |
| Outcome: | The proposed model improves performance on several temporal reasoning benchmarks and establishes new state-of-the-art results. |
CARVQ: Corrective Adaptor with Group Residual Vector Quantization for LLM Embedding Compression (2025.findings-emnlp)
Copied to clipboard
Dayin Gou, Sanghyun Byun, Nilesh Malpeddi, Gabrielle De Micheli, Prathamesh Vaste, Jacob Song, Woo Seong Chung
| Challenge: | Large Language Models typically rely on a large number of parameters for token embedding, leading to substantial storage requirements and memory footprints. |
| Approach: | They propose a corrective Adaptor with group Residual Vector Quantization that can be used to compress the embedding layer without requiring specialized hardware. |
| Outcome: | The proposed corrective adaptor can achieve lower average bitwidth-per-parameter while maintaining reasonable perplexity and accuracy compared to scalar quantization. |
Language Models Can be Efficiently Steered via Minimal Embedding Layer Transformations (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for fine-tuning Large Language Models (LLMs) neglect the embedding layer. |
| Approach: | They propose a PEFT approach that modifies input embeddings without altering hidden layers. |
| Outcome: | Experiments show that TinyTE modifies embeddings without altering hidden layers . the proposed approach achieves competitive performance while requiring 0.0001% of parameters . |